# Snow CLI User Guide - Codebase Setup

Welcome to Snow CLI! Agentic coding in your terminal.

## Codebase Setup

Snow CLI supports enabling local codebase functionality.

_The codebase is a vector search-based SQLite database used to store source code and comments from your codebase, with natural language query capabilities through vectorization._

## Configuration Storage

Codebase configuration is split into two parts:

- **Project-level config** (`.snow/codebase.json`): Stored in project root, controls enable/disable status, indexing parameters, reranking config, etc.
- **Global config** (`~/.snow/codebase.json`): Stores Embedding service configuration, shared across projects

This allows each project to independently control codebase functionality, while Embedding settings only need to be configured once.

## Quick Toggle

Use the `/codebase` command to quickly control codebase functionality for the current project:

- `/codebase` - Toggle enable/disable
- `/codebase on` - Enable codebase
- `/codebase off` - Disable codebase
- `/codebase status` - View current status

When enabling for the first time, you need to configure the Embedding service in `/home` first.

## Configuration UI

In `/home` → Codebase Config, settings are organized into collapsible groups to save screen space:

```
  CodeBase Enabled:            ← Master toggle
  Agent Review:                ← AI review of search results (mutually exclusive with reranking)
  Result Reranking:            ← Rerank search results (mutually exclusive with agent review)
  ▶ Embedding Model Config     ← Press Enter to expand/collapse
  ▶ Reranking Model Config     ← Press Enter to expand/collapse
  ▶ Batch Settings             ← Press Enter to expand/collapse
```

Use ↑↓ to navigate, Enter to edit/toggle/expand, Ctrl+S or Esc to save.

## Search Result Optimization

After codebase search returns results, there are two optimization modes available (**mutually exclusive, cannot be enabled simultaneously**):

### Agent Review

Uses an AI model (basicModel) to semantically review search results, filtering out irrelevant items and potentially suggesting better search keywords. Best for scenarios requiring deep code semantic understanding.

- Supports multi-round retry with keyword suggestions
- Can identify high-confidence files for deep exploration
- Depends on configured AI models (basicModel / advancedModel)

### Result Reranking

Uses a dedicated Rerank model to reorder search results by relevance, returning the Top N most relevant items. More lightweight and efficient compared to Agent Review, best for speed-oriented scenarios.

- Calls standard Rerank API (compatible with Jina Reranker, Cohere Rerank, etc.)
- Built-in 3-attempt retry with exponential backoff
- Built-in context length protection: uses tiktoken for precise token counting, auto-truncates or drops oversized documents to prevent context overflow
- Graceful degradation to raw search results on failure

**Mutual exclusivity**: Enabling "Result Reranking" automatically disables "Agent Review", and vice versa. The reranking model must be configured before it can be enabled.

## Embedding Service Configuration

Expand under "▶ Embedding Model Config":

- The codebase supports three request schemes: Jina (OpenAI-compatible), Ollama (local deployment, supports both OpenAI-compatible `/v1/embeddings` and native `/api/embed`), and Gemini.

- Codebase BaseURL (multiple forms supported; Snow CLI will auto-normalize to the final endpoint):

  - Jina (OpenAI-compatible) supports: `https://api.jina.ai`, `https://api.jina.ai/v1`, `https://api.jina.ai/v1/embeddings` (final request: `.../v1/embeddings`).
  - Ollama supports: `http://localhost:11434`, `http://localhost:11434/v1`, `http://localhost:11434/v1/embeddings` (OpenAI-compatible); and `http://localhost:11434/api`, `http://localhost:11434/api/embed` (Ollama native).

- Embedding Dimensions: Enter the dimensions supported by your embedding model. Some providers may ignore the `dimensions` parameter; Snow CLI will log a warning if the returned dimensions don't match.

## Reranking Model Configuration

Expand under "▶ Reranking Model Config":

| Setting | Description | Default |
|---------|-------------|---------|
| Model Name | Rerank model name, e.g. `jina-reranker-v2-base-multilingual` | — |
| Base URL | Rerank API endpoint, e.g. `https://api.jina.ai` (auto-appends `/v1/rerank`) | — |
| API Key | API authentication key (optional, can be left empty for local deployments) | — |
| Model Context Length | Maximum context tokens the model supports, used to prevent request overflow | 4096 |
| Top N | Number of top results to return after reranking | 5 |

**Context length protection**: Before sending requests, tiktoken is used to precisely calculate the total token count of all documents. Individual documents exceeding 30% of the context window are truncated; documents that would exceed the total budget are dropped. This ensures requests never overflow the model's context limit.

## Indexing Parameters

Expand under "▶ Batch Settings":

- Chunking Configuration: Configure how code is split into chunks for indexing. These settings control the size and overlap of code segments:
  - `maxLinesPerChunk`: Maximum lines per chunk (default: 200)
  - `minLinesPerChunk`: Minimum lines per chunk (default: 10)
  - `minCharsPerChunk`: Minimum characters per chunk (default: 20)
  - `overlapLines`: Lines overlapping between consecutive chunks (default: 20)
    These settings affect search accuracy and indexing performance.
- Batch Processing: Control how files are processed in batches for efficient indexing:
  - `maxLines`: Maximum lines per batch request (default: 10)
  - `concurrency`: Number of concurrent batch operations (default: 3)
    This controls the number of `input` items sent to the embedding API per request.

**Note: Maximum batch lines refers to the number of `input` items in the request body, not the number of lines in code slices**

## Related Features

After enabling codebase indexing, the following features will be significantly enhanced:

- [Vulnerability Hunting Mode](./11.Vulnerability%20Hunting%20Mode.md) - Codebase indexing can greatly improve accuracy and efficiency of security analysis
- [Command Panel Guide](./09.0.Command%20Panel%20Guide.md) - Use `/reindex` to rebuild codebase index, use `/codebase` to toggle enable/disable
